EconBase
← Back to paper

Regression Discontinuity Design with Potentially Many Covariates

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

10,139,119 characters

Regression Discontinuity Design with Potentially Many Covariates


\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\multicolumn{2}{c}{{\small{}Estimation methods:}} & \multicolumn{2}{c}{{\small{}Standard}} & \multicolumn{2}{c}{{\small{}Covariate adjusted}} & \multicolumn{2}{c}{{\small{}Covariate adjusted}} & \multicolumn{2}{c}{{\small{}Covariate selection}}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}$p$} & {\small{}Bias} & {\small{}RMSE} & {\small{}Bias} & {\small{}RMSE} & {\small{}Bias} & {\small{}RMSE} & {\small{}Bias} & {\small{}RMSE}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP1} & {\small{}5} & {\small{}0.022} & {\small{}0.062} & {\small{}0.023} & {\small{}0.064} & {\small{}0.022} & {\small{}0.065} & {\small{}0.022} & {\small{}0.062}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.021} & {\small{}0.065} & {\small{}0.022} & {\small{}0.068} & {\small{}0.019} & {\small{}0.070} & {\small{}0.020} & {\small{}0.065}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.020} & {\small{}0.061} & {\small{}0.019} & {\small{}0.066} & {\small{}0.013} & {\small{}0.074} & {\small{}0.019} & {\small{}0.061}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.021} & {\small{}0.064} & {\small{}0.022} & {\small{}0.074} & {\small{}0.018} & {\small{}0.089} & {\small{}0.020} & {\small{}0.064}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.022} & {\small{}0.064} & {\small{}0.022} & {\small{}0.080} & {\small{}0.002} & {\small{}0.323} & {\small{}0.022} & {\small{}0.064}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.021} & {\small{}0.065} & {\small{}0.027} & {\small{}0.101} & {\small{}0.003} & {\small{}0.316} & {\small{}0.021} & {\small{}0.065}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.019} & {\small{}0.063} & {\small{}0.038} & {\small{}0.387} & {\small{}0.004} & {\small{}0.133} & {\small{}0.018} & {\small{}0.063}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.018} & {\small{}0.065} & {\small{}0.026} & {\small{}0.082} & {\small{}0.006} & {\small{}0.102} & {\small{}0.017} & {\small{}0.065}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.021} & {\small{}0.064} & {\small{}0.028} & {\small{}0.072} & {\small{}0.009} & {\small{}0.100} & {\small{}0.020} & {\small{}0.064}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP2} & {\small{}5} & {\small{}0.019} & {\small{}0.540} & {\small{}0.000} & {\small{}0.058} & {\small{}0.029} & {\small{}0.081} & {\small{}0.029} & {\small{}0.090}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}-0.007} & {\small{}0.525} & {\small{}-0.001} & {\small{}0.061} & {\small{}0.026} & {\small{}0.084} & {\small{}0.028} & {\small{}0.094}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.008} & {\small{}0.543} & {\small{}-0.007} & {\small{}0.064} & {\small{}0.020} & {\small{}0.094} & {\small{}0.026} & {\small{}0.099}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}-0.006} & {\small{}0.530} & {\small{}-0.006} & {\small{}0.070} & {\small{}0.025} & {\small{}0.123} & {\small{}0.027} & {\small{}0.105}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.002} & {\small{}0.521} & {\small{}-0.012} & {\small{}0.080} & {\small{}-0.001} & {\small{}1.485} & {\small{}0.026} & {\small{}0.102}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.043} & {\small{}0.531} & {\small{}-0.010} & {\small{}0.098} & {\small{}0.002} & {\small{}1.116} & {\small{}0.025} & {\small{}0.105}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.042} & {\small{}0.544} & {\small{}-0.033} & {\small{}0.551} & {\small{}-0.052} & {\small{}0.874} & {\small{}0.023} & {\small{}0.107}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}-0.007} & {\small{}0.505} & {\small{}-0.059} & {\small{}1.162} & {\small{}0.027} & {\small{}1.166} & {\small{}0.019} & {\small{}0.116}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.003} & {\small{}0.535} & {\small{}-0.058} & {\small{}1.636} & {\small{}-0.066} & {\small{}1.578} & {\small{}0.027} & {\small{}0.118}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP3} & {\small{}5} & {\small{}0.013} & {\small{}0.708} & {\small{}-0.004} & {\small{}0.058} & {\small{}0.029} & {\small{}0.081} & {\small{}0.028} & {\small{}0.103}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}-0.010} & {\small{}0.681} & {\small{}-0.005} & {\small{}0.060} & {\small{}0.026} & {\small{}0.084} & {\small{}0.030} & {\small{}0.179}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.002} & {\small{}0.706} & {\small{}-0.011} & {\small{}0.064} & {\small{}0.020} & {\small{}0.094} & {\small{}0.005} & {\small{}0.177}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}-0.009} & {\small{}0.691} & {\small{}-0.011} & {\small{}0.070} & {\small{}0.025} & {\small{}0.123} & {\small{}0.030} & {\small{}0.186}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.000} & {\small{}0.683} & {\small{}-0.017} & {\small{}0.080} & {\small{}0.021} & {\small{}0.321} & {\small{}0.018} & {\small{}0.176}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.038} & {\small{}0.686} & {\small{}-0.015} & {\small{}0.085} & {\small{}0.042} & {\small{}0.784} & {\small{}0.017} & {\small{}0.190}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.048} & {\small{}0.722} & {\small{}-0.036} & {\small{}0.301} & {\small{}0.045} & {\small{}0.605} & {\small{}0.013} & {\small{}0.201}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}-0.019} & {\small{}0.668} & {\small{}-0.055} & {\small{}0.347} & {\small{}-0.021} & {\small{}0.957} & {\small{}0.009} & {\small{}0.216}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}-0.004} & {\small{}0.707} & {\small{}-0.050} & {\small{}0.485} & {\small{}0.015} & {\small{}0.965} & {\small{}0.014} & 0.217\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP4} & {\small{}5} & {\small{}0.028} & {\small{}0.180} & {\small{}0.018} & {\small{}0.066} & {\small{}0.029} & {\small{}0.081} & {\small{}0.022} & {\small{}0.131}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.021} & {\small{}0.187} & {\small{}0.015} & {\small{}0.068} & {\small{}0.026} & {\small{}0.084} & {\small{}0.019} & {\small{}0.185}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.019} & {\small{}0.193} & {\small{}0.010} & {\small{}0.071} & {\small{}0.020} & {\small{}0.094} & {\small{}0.015} & {\small{}0.192}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.028} & {\small{}0.197} & {\small{}0.013} & {\small{}0.080} & {\small{}0.025} & {\small{}0.123} & {\small{}0.024} & {\small{}0.196}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.021} & {\small{}0.192} & {\small{}0.008} & {\small{}0.092} & {\small{}0.001} & {\small{}1.480} & {\small{}0.019} & {\small{}0.190}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.029} & {\small{}0.198} & {\small{}0.011} & {\small{}0.110} & {\small{}-0.026} & {\small{}1.491} & {\small{}0.022} & {\small{}0.194}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.036} & {\small{}0.206} & {\small{}-0.053} & {\small{}0.853} & {\small{}-0.024} & {\small{}0.470} & {\small{}0.032} & {\small{}0.203}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.017} & {\small{}0.194} & {\small{}-0.039} & {\small{}0.547} & {\small{}-0.072} & {\small{}0.578} & {\small{}0.013} & {\small{}0.190}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.020} & {\small{}0.196} & {\small{}-0.026} & {\small{}0.549} & {\small{}-0.053} & {\small{}0.627} & {\small{}0.016} & {\small{}0.194}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\small\par}
\end{table}

\newpage{}

\begin{table}[htp]
{\small{}\caption{{\small{}\label{tab:selected}}Simulation: Number of selected covariates}
}{\small\par}
\centering{}{\small{}}
\begin{tabular}{lcccc}
\hline
 & {\small{}$p$} & {\small{}Average} & {\small{}Min} & {\small{}Max}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP1} & {\small{}5} & {\small{}0.376} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.362} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.357} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.353} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.332} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.379} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.350} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.331} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.264} & {\small{}0} & {\small{}1}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP2} & {\small{}5} & {\small{}2.868} & {\small{}2} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}2.780} & {\small{}2} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}2.689} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}2.631} & {\small{}2} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}2.555} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}2.574} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}2.478} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}2.341} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}2.303} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP3} & {\small{}5} & {\small{}3.910} & {\small{}3} & {\small{}5}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}3.063} & {\small{}2} & {\small{}5}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}3.045} & {\small{}2} & {\small{}5}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}3.003} & {\small{}1} & {\small{}4}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}2.959} & {\small{}1} & {\small{}5}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}2.920} & {\small{}1} & {\small{}5}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}2.850} & {\small{}1} & {\small{}4}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}2.683} & {\small{}1} & {\small{}4}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}2.547} & {\small{}1} & {\small{}4}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP4} & {\small{}5} & {\small{}1.771} & {\small{}1} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.983} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.860} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.787} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.719} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.730} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.680} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.593} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.479} & {\small{}0} & {\small{}3}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\small\par}
\end{table}

\newpage{}

\begin{table}[htp]
{\small{}\caption{{\small{}\label{tab:inf}}Simulation: Inference}
}{\small\par}
\centering{}{\small{}}
\begin{tabular}{lccccccccc}
\hline
\multicolumn{2}{c}{{\small{}MSE-optimal bandwidths:}} & \multicolumn{4}{c}{{\small{}w/o Covariates}} & \multicolumn{2}{c}{{\small{}w/ Covariates}} & \multicolumn{2}{c}{{\small{}Adaptive}}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\multicolumn{2}{c}{{\small{}Estimation methods:}} & \multicolumn{2}{c}{{\small{}Standard}} & \multicolumn{2}{c}{{\small{}Covariate adjusted}} & \multicolumn{2}{c}{{\small{}Covariate adjusted}} & \multicolumn{2}{c}{{\small{}Covariate selection}}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}$p$} & {\small{}CP} & {\small{}Length} & {\small{}CP} & {\small{}Length} & {\small{}CP} & {\small{}Length} & {\small{}CP} & {\small{}Length}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP1} & {\small{}5} & {\small{}0.918} & {\small{}0.270} & {\small{}0.898} & {\small{}0.246} & {\small{}0.901} & {\small{}0.246} & {\small{}0.916} & {\small{}0.270}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.892} & {\small{}0.213} & {\small{}0.855} & {\small{}0.211} & {\small{}0.856} & {\small{}0.205} & {\small{}0.891} & {\small{}0.213}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.925} & {\small{}0.293} & {\small{}0.815} & {\small{}0.259} & {\small{}0.779} & {\small{}0.280} & {\small{}0.919} & {\small{}0.275}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.914} & {\small{}0.260} & {\small{}0.734} & {\small{}0.148} & {\small{}0.658} & {\small{}0.156} & {\small{}0.914} & {\small{}0.260}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.915} & {\small{}0.241} & {\small{}0.649} & {\small{}0.182} & {\small{}0.476} & {\small{}0.107} & {\small{}0.916} & {\small{}0.241}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.915} & {\small{}0.246} & {\small{}0.540} & {\small{}0.152} & {\small{}0.247} & {\small{}0.076} & {\small{}0.917} & {\small{}0.246}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.913} & {\small{}0.246} & {\small{}0.194} & {\small{}0.342} & {\small{}0.186} & {\small{}0.040} & {\small{}0.911} & {\small{}0.246}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.900} & {\small{}0.252} & {\small{}0.108} & {\small{}0.040} & {\small{}0.174} & {\small{}0.048} & {\small{}0.904} & {\small{}0.233}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.905} & {\small{}0.188} & {\small{}0.138} & {\small{}0.020} & {\small{}0.155} & {\small{}0.065} & {\small{}0.903} & {\small{}0.188}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP2} & {\small{}5} & {\small{}0.926} & {\small{}2.474} & {\small{}0.661} & {\small{}0.232} & {\small{}0.845} & {\small{}0.283} & {\small{}0.867} & {\small{}0.331}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.932} & {\small{}1.732} & {\small{}0.627} & {\small{}0.209} & {\small{}0.781} & {\small{}0.292} & {\small{}0.873} & {\small{}0.413}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.922} & {\small{}2.261} & {\small{}0.591} & {\small{}0.261} & {\small{}0.677} & {\small{}0.252} & {\small{}0.876} & {\small{}0.395}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.938} & {\small{}1.932} & {\small{}0.517} & {\small{}0.136} & {\small{}0.505} & {\small{}0.143} & {\small{}0.864} & {\small{}0.305}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.934} & {\small{}2.351} & {\small{}0.472} & {\small{}0.170} & {\small{}0.305} & {\small{}0.122} & {\small{}0.879} & {\small{}0.344}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.930} & {\small{}1.853} & {\small{}0.407} & {\small{}0.127} & {\small{}0.267} & {\small{}0.231} & {\small{}0.884} & {\small{}0.300}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.918} & {\small{}2.063} & {\small{}0.250} & {\small{}0.069} & {\small{}0.566} & {\small{}2.065} & {\small{}0.876} & {\small{}0.570}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.932} & {\small{}2.538} & {\small{}0.778} & {\small{}2.501} & {\small{}0.808} & {\small{}3.648} & {\small{}0.867} & {\small{}0.559}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.915} & {\small{}2.286} & {\small{}0.881} & {\small{}6.466} & {\small{}0.875} & {\small{}5.146} & {\small{}0.876} & {\small{}0.494}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP3} & {\small{}5} & {\small{}0.926} & {\small{}3.173} & {\small{}0.615} & {\small{}0.219} & {\small{}0.845} & {\small{}0.283} & {\small{}0.870} & {\small{}0.456}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.932} & {\small{}2.243} & {\small{}0.601} & {\small{}0.202} & {\small{}0.781} & {\small{}0.292} & {\small{}0.905} & {\small{}0.724}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.922} & {\small{}2.593} & {\small{}0.564} & {\small{}0.236} & {\small{}0.677} & {\small{}0.252} & {\small{}0.908} & {\small{}0.609}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.931} & {\small{}2.498} & {\small{}0.495} & {\small{}0.151} & {\small{}0.505} & {\small{}0.143} & {\small{}0.881} & {\small{}0.680}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.928} & {\small{}3.277} & {\small{}0.456} & {\small{}0.174} & {\small{}0.295} & {\small{}0.122} & {\small{}0.914} & {\small{}1.118}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.927} & {\small{}2.462} & {\small{}0.397} & {\small{}0.132} & {\small{}0.189} & {\small{}0.082} & {\small{}0.888} & {\small{}0.444}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.912} & {\small{}2.750} & {\small{}0.182} & {\small{}0.073} & {\small{}0.164} & {\small{}0.116} & {\small{}0.898} & {\small{}0.846}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.931} & {\small{}3.578} & {\small{}0.119} & {\small{}0.139} & {\small{}0.175} & {\small{}0.362} & {\small{}0.895} & {\small{}0.745}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.922} & {\small{}3.235} & {\small{}0.130} & {\small{}0.158} & {\small{}0.157} & {\small{}0.255} & {\small{}0.892} & {\small{}0.635}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP4} & {\small{}5} & {\small{}0.910} & {\small{}1.083} & {\small{}0.807} & {\small{}0.269} & {\small{}0.845} & {\small{}0.283} & {\small{}0.899} & {\small{}0.525}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.924} & {\small{}0.619} & {\small{}0.762} & {\small{}0.253} & {\small{}0.781} & {\small{}0.292} & {\small{}0.912} & {\small{}0.624}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.923} & {\small{}0.948} & {\small{}0.689} & {\small{}0.277} & {\small{}0.677} & {\small{}0.252} & {\small{}0.904} & {\small{}0.854}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.912} & {\small{}0.631} & {\small{}0.588} & {\small{}0.149} & {\small{}0.505} & {\small{}0.143} & {\small{}0.909} & {\small{}0.631}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.932} & {\small{}0.993} & {\small{}0.489} & {\small{}0.162} & {\small{}0.296} & {\small{}0.122} & {\small{}0.925} & {\small{}0.997}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.920} & {\small{}0.666} & {\small{}0.375} & {\small{}0.113} & {\small{}0.232} & {\small{}0.121} & {\small{}0.914} & {\small{}0.673}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.914} & {\small{}0.714} & {\small{}0.306} & {\small{}0.289} & {\small{}0.471} & {\small{}0.835} & {\small{}0.914} & {\small{}0.714}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.916} & {\small{}0.918} & {\small{}0.677} & {\small{}1.547} & {\small{}0.705} & {\small{}1.436} & {\small{}0.920} & {\small{}0.882}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.910} & {\small{}0.821} & {\small{}0.853} & {\small{}1.862} & {\small{}0.791} & {\small{}3.522} & {\small{}0.910} & {\small{}0.810}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\small\par}
\end{table}

\newpage{}

\begin{table}[htp]
{\small{}\caption{{\small{}\label{tab:band}}Simulation: MSE-optimal bandwidths}
}{\small\par}
\centering{}{\small{}}
\begin{tabular}{lccccccc}
\hline
\multicolumn{2}{c}{{\small{}MSE-optimal bandwidths:}} & \multicolumn{2}{c}{{\small{}w/o Covariates}} & \multicolumn{2}{c}{{\small{}w/ Covariates}} & \multicolumn{2}{c}{{\small{}Adaptive}}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}$p$} & {\small{}Mean} & {\small{}SD} & {\small{}Mean} & {\small{}SD} & {\small{}Mean} & {\small{}SD}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP1} & {\small{}5} & {\small{}0.198} & {\small{}0.046} & {\small{}0.190} & {\small{}0.043} & {\small{}0.197} & {\small{}0.046}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.195} & {\small{}0.046} & {\small{}0.179} & {\small{}0.041} & {\small{}0.195} & {\small{}0.046}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.197} & {\small{}0.043} & {\small{}0.163} & {\small{}0.035} & {\small{}0.197} & {\small{}0.044}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.196} & {\small{}0.044} & {\small{}0.147} & {\small{}0.031} & {\small{}0.196} & {\small{}0.045}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.197} & {\small{}0.044} & {\small{}0.128} & {\small{}0.026} & {\small{}0.196} & {\small{}0.043}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.194} & {\small{}0.045} & {\small{}0.108} & {\small{}0.024} & {\small{}0.193} & {\small{}0.045}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.196} & {\small{}0.047} & {\small{}0.075} & {\small{}0.024} & {\small{}0.196} & {\small{}0.047}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.194} & {\small{}0.045} & {\small{}0.069} & {\small{}0.019} & {\small{}0.194} & {\small{}0.045}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.197} & {\small{}0.043} & {\small{}0.071} & {\small{}0.020} & {\small{}0.196} & {\small{}0.043}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP2} & {\small{}5} & {\small{}0.177} & {\small{}0.025} & {\small{}0.108} & {\small{}0.011} & {\small{}0.111} & {\small{}0.013}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.177} & {\small{}0.025} & {\small{}0.107} & {\small{}0.011} & {\small{}0.113} & {\small{}0.014}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.176} & {\small{}0.026} & {\small{}0.105} & {\small{}0.011} & {\small{}0.115} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.177} & {\small{}0.025} & {\small{}0.102} & {\small{}0.011} & {\small{}0.115} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.176} & {\small{}0.025} & {\small{}0.099} & {\small{}0.012} & {\small{}0.117} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.176} & {\small{}0.026} & {\small{}0.094} & {\small{}0.014} & {\small{}0.116} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.176} & {\small{}0.025} & {\small{}0.130} & {\small{}0.026} & {\small{}0.118} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.176} & {\small{}0.025} & {\small{}0.159} & {\small{}0.029} & {\small{}0.121} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.177} & {\small{}0.025} & {\small{}0.173} & {\small{}0.031} & {\small{}0.121} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP3} & {\small{}5} & {\small{}0.182} & {\small{}0.028} & {\small{}0.108} & {\small{}0.011} & {\small{}0.117} & {\small{}0.014}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.183} & {\small{}0.029} & {\small{}0.107} & {\small{}0.011} & {\small{}0.138} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.181} & {\small{}0.029} & {\small{}0.105} & {\small{}0.011} & {\small{}0.139} & {\small{}0.017}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.183} & {\small{}0.028} & {\small{}0.102} & {\small{}0.011} & {\small{}0.138} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.182} & {\small{}0.028} & {\small{}0.099} & {\small{}0.012} & {\small{}0.139} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.183} & {\small{}0.029} & {\small{}0.092} & {\small{}0.013} & {\small{}0.140} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.183} & {\small{}0.029} & {\small{}0.078} & {\small{}0.021} & {\small{}0.141} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.181} & {\small{}0.028} & {\small{}0.076} & {\small{}0.022} & {\small{}0.142} & {\small{}0.018}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.183} & {\small{}0.029} & {\small{}0.074} & {\small{}0.021} & {\small{}0.143} & {\small{}0.017}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\small{}DGP4} & {\small{}5} & {\small{}0.141} & {\small{}0.015} & {\small{}0.108} & {\small{}0.011} & {\small{}0.126} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}10} & {\small{}0.144} & {\small{}0.015} & {\small{}0.107} & {\small{}0.011} & {\small{}0.142} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}20} & {\small{}0.144} & {\small{}0.016} & {\small{}0.105} & {\small{}0.011} & {\small{}0.142} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}30} & {\small{}0.144} & {\small{}0.015} & {\small{}0.102} & {\small{}0.011} & {\small{}0.143} & {\small{}0.015}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}40} & {\small{}0.145} & {\small{}0.015} & {\small{}0.099} & {\small{}0.012} & {\small{}0.144} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}50} & {\small{}0.144} & {\small{}0.016} & {\small{}0.092} & {\small{}0.013} & {\small{}0.143} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}100} & {\small{}0.145} & {\small{}0.015} & {\small{}0.104} & {\small{}0.021} & {\small{}0.143} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}250} & {\small{}0.144} & {\small{}0.016} & {\small{}0.132} & {\small{}0.025} & {\small{}0.143} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\small{}500} & {\small{}0.144} & {\small{}0.015} & {\small{}0.150} & {\small{}0.025} & {\small{}0.143} & {\small{}0.016}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\small\par}
\end{table}

\newpage{}

\section{Empirical illustration: Head Start data\label{sec:emp}}

To illustrate our variable selection approach, we revisit the problem
of the Head Start program first studied by Ludwig and Miller (2007)
where they investigate the effect of the Head Start program on various
outcomes related to health and schooling. The federal government provided
grant-writing assistance to the 300 poorest counties based on the
poverty index to apply for the Head Start program. This leads to the
RDD with the poverty index as a running variable where the cut-off
value is set as $\bar{x}=59.1984$. Ludwig and Miller (2007) conducted
their RDD analysis using no covariate, and CCFT examined the impact
of the covariance-adjustment. CCFT employed nine pre-intervention
covariates from the U.S. Census, which include total population, percentage
of population, percentages of black and urban population, and levels
and percentages of population in three age groups (children aged 3
to 5, children aged 14 to 17, and adults older than 25). The main
finding by CCFT is that the covariate adjusted RDD inference yields
shorter confidence intervals while the RDD point estimates remain
stable.

An important aspect of the Head start example is that it is unclear
which covariates become useful to improve efficiency mainly due to
the lack of economic theories behind the problem. We conduct the empirical
exercises of CCFT by applying our variable selection approach with
two extensions. First, we introduce 36 interaction terms in addition
to the nine original covariates. Second, we also implement those estimation
and inference for subsamples to see the effect of changes in the ratio
of the number of covariates ($p$) to that of observations ($n$).
Hereafter, as in CCFT, we focus on child mortality among many outcome
variables.

Table 5 shows the results of our empirical illustration. Four columns
correspond to four estimation procedures which are the same as those
used in the simulation experiments. The first panel shows the full
sample results ($n=2799$ and $p/n=0.016$). The RDD causal effect
estimates are presented in the first row. The next three rows show
95\% confidence intervals, their percentage length changes relative
to the one in the first column, and their associated $p$-values where
these are obtained without restriction on the MSE optimal bandwidth
for the local linear regression ($h$) and the pilot bandwidth ($b$).
See CCT and CCFT for more details on the robust inference methods.
These results are also obtained under the restriction $h/b=1$, which
are reported in the following three rows. The next two rows in the
same panel present the bandwidths $(h,b)$, and effective sample sizes
$(n_{-},n_{+})$ used for the RDD estimation. The effective sample
sizes are the numbers of observations of the running variable in the
intervals $[\bar{x}-h,\bar{x}]$ and $[\bar{x},\bar{x}+h]$. We also
report the selected covariates for our covariate selection approach.
We use subsamples of the first 1000 and 500 observations for the second
and third panels, leading to $p/n=0.045$ and $0.090$, respectively.

For the full sample case, the covariate adjusted estimates mildly
deviate from the standard one while our estimate based on the variable
selection is identical to the standard one. Although the confidence
intervals of the covariate adjusted approaches are shorter than the
standard one, this might induce under-coverages for the case of many
covariates as illustrated in the simulation experiment. As the sample
size gets smaller, the observations made here are amplified. In contrast,
we can see the stable performance of the variable selection approach
and its mild contribution to shorten the confidence intervals. Table
6 presents the corresponding results for the case of nine covariates
as considered by CCFT. The results are essentially similar to those
of Table 5, although the influence of the covariates is less dramatic.

\newpage{}

{\footnotesize{}}
\begin{table}[H]
{\footnotesize{}\caption{Empirical illustration: Head Start data ($45$ covariates)}
}{\footnotesize\par}
\centering{}{\footnotesize{}}
\begin{tabular}{llcccc}
\hline
 & {\footnotesize{}MSE-optimal bandwidths:} & \multicolumn{2}{c}{{\footnotesize{}w/o Covariates}} & {\footnotesize{}w/ Covariates} & {\footnotesize{}Adaptive}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
 & {\footnotesize{}Estimation methods:} & {\footnotesize{}Standard} & {\footnotesize{}Cov-adjusted} & {\footnotesize{}Cov-adjusted} & {\footnotesize{}Variable selection}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=2779$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-2.41$} & {\footnotesize{}$-1.82$} & {\footnotesize{}$-3.64$} & {\footnotesize{}$-2.41$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.016$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-5.46,-0.1]$} & {\footnotesize{}$[-3.85,-0.07]$} & {\footnotesize{}$[-6.05,-1.24]$} & {\footnotesize{}$[-5.46,-0.1]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-29.57$} & {\footnotesize{}$-10.31$} & {\footnotesize{}$0$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}$0.042$} & {\footnotesize{}$0.042$} & {\footnotesize{}$0.003$} & {\footnotesize{}$0.042$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-6.41,-1.09]$} & {\footnotesize{}$[-5.44,-1.10]$} & {\footnotesize{}$[-6.55,-1.14]$} & {\footnotesize{}$[-6.41,-1.09]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-18.37$} & {\footnotesize{}$1.64$} & {\footnotesize{}$0$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.006} & {\footnotesize{}0.003} & {\footnotesize{}0.005} & {\footnotesize{}0.006}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.81, 10.73} & {\footnotesize{}6.81, 10.73} & {\footnotesize{}3.02, 5.40} & {\footnotesize{}6.81, 10.73}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}234, 180} & {\footnotesize{}234, 180} & {\footnotesize{}96, 84} & {\footnotesize{}234, 180}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}None}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=1000$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-1.68$} & {\footnotesize{}$-2.40$} & {\footnotesize{}$-3.44$} & {\footnotesize{}$-1.48$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.045$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-5.45,1.75]$} & {\footnotesize{}$[-5.44,-0.29]$} & {\footnotesize{}$[-6.14,-1.34]$} & {\footnotesize{}$[-5.08,1.79]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-28.38$} & {\footnotesize{}$-33.23$} & {\footnotesize{}$-4.49$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.314} & {\footnotesize{}0.029} & {\footnotesize{}0.002} & {\footnotesize{}0.347}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-8.26,0.22]$} & {\footnotesize{}$[-6.84,-0.47]$} & {\footnotesize{}$[-5.46,.0.15]$} & {\footnotesize{}$[-7.94,0.27]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-24.95$} & {\footnotesize{}$-33.87$} & {\footnotesize{}$-3.28$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.063} & {\footnotesize{}0.025} & {\footnotesize{}0.064} & {\footnotesize{}0.070}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.52, 10.23} & {\footnotesize{}6.52, 10.23} & {\footnotesize{}3.47, 6.00} & {\footnotesize{}5.26, 8.07}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}74, 77} & {\footnotesize{}79, 79} & {\footnotesize{}40, 46} & {\footnotesize{}79, 79}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}\% of adult population}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=500$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-2.35$} & {\footnotesize{}$-4.23$} & {\footnotesize{}$-5.68$} & {\footnotesize{}$-2.22$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.090$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-7.25,2.48]$} & {\footnotesize{}$[-7.44,-2.06]$} & {\footnotesize{}$[-9.64,-4.63]$} & {\footnotesize{}$[-6.93,2.25]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-44.77$} & {\footnotesize{}$-48.58$} & {\footnotesize{}$-5.74$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.337} & {\footnotesize{}0.001} & {\footnotesize{}0.000} & {\footnotesize{}0.317}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-10.22,1.42]$} & {\footnotesize{}$[-9.03,-3.43]$} & {\footnotesize{}$[-9.62,-4.29]$} & {\footnotesize{}$[-10,1.07]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-51.93$} & {\footnotesize{}$-54.16$} & {\footnotesize{}$-4.94$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.139} & {\footnotesize{}0.000} & {\footnotesize{}0.000} & {\footnotesize{}0.110}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.37, 9.16} & {\footnotesize{}6.37, 9.16} & {\footnotesize{}4.96, 7.49} & {\footnotesize{}4.31, 6.82}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}60, 56} & {\footnotesize{}61, 56} & {\footnotesize{}49, 47} & {\footnotesize{}61, 56}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}\% of adult population}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\footnotesize\par}
\end{table}
{\footnotesize\par}

{\footnotesize{}}
\begin{table}[H]
{\footnotesize{}\caption{Empirical illustration: Head Start data ($9$ covariates)}
}{\footnotesize\par}
\centering{}{\footnotesize{}}
\begin{tabular}{llcccc}
\hline
 & {\footnotesize{}MSE-optimal bandwidths:} & \multicolumn{2}{c}{{\footnotesize{}w/o Covariates}} & {\footnotesize{}w/ Covariates} & {\footnotesize{}Adaptive}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
 & {\footnotesize{}Estimation methods:} & {\footnotesize{}Standard} & {\footnotesize{}Cov-adjusted} & {\footnotesize{}Cov-adjusted} & {\footnotesize{}Variable selection}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=2779$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-2.41$} & {\footnotesize{}$-2.51$} & {\footnotesize{}$2.47$} & {\footnotesize{}-2.41}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.003$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-5.46,-0.1]$} & {\footnotesize{}$[-5.37,-0.45]$} & {\footnotesize{}$[-5.21,-0.37]$} & {\footnotesize{}$[-5.46,-0.1]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-8.24$} & {\footnotesize{}$-9.74$} & {\footnotesize{}$0$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}$0.042$} & {\footnotesize{}$0.021$} & {\footnotesize{}$0.024$} & {\footnotesize{}$0.042$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-6.41,-1.09]$} & {\footnotesize{}$[-6.63,-1.46]$} & {\footnotesize{}$[-6.54,-1.39]$} & {\footnotesize{}$[-6.41,-1.09]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-2.87$} & {\footnotesize{}$-3.23$} & {\footnotesize{}$0$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.006} & {\footnotesize{}0.002} & {\footnotesize{}0.003} & {\footnotesize{}0.006}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.81, 10.73} & {\footnotesize{}6.81, 10.73} & {\footnotesize{}6.98, 11.64} & {\footnotesize{}6.81, 10.73}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}234, 180} & {\footnotesize{}234, 180} & {\footnotesize{}240, 184} & {\footnotesize{}234, 180}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}None}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=1000$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-1.68$} & {\footnotesize{}$-2.05$} & {\footnotesize{}$-1.53$} & {\footnotesize{}$-1.48$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.009$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-5.45,1.75]$} & {\footnotesize{}$[-6.35,-1.1]$} & {\footnotesize{}$[-7.7,-2.28]$} & {\footnotesize{}$[-5.08,1.79]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-7.12$} & {\footnotesize{}$-20.0$} & {\footnotesize{}$-4.49$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.314} & {\footnotesize{}0.141} & {\footnotesize{}0.237} & {\footnotesize{}0.347}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-8.26,0.22]$} & {\footnotesize{}$[-8.26,-0.63]$} & {\footnotesize{}$[-5.93,-0.88]$} & {\footnotesize{}$[-7.94,0.27]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-10.20$} & {\footnotesize{}$-19.75$} & {\footnotesize{}$-3.28$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.063} & {\footnotesize{}0.022} & {\footnotesize{}0.145} & {\footnotesize{}0.070}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.52, 10.23} & {\footnotesize{}6.52, 10.23} & {\footnotesize{}9.14, 13.93} & {\footnotesize{}9.26, 8.07}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}74, 77} & {\footnotesize{}79, 79} & {\footnotesize{}121, 98} & {\footnotesize{}79, 79}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}\% of adult population}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
{\footnotesize{}$n=500$} & {\footnotesize{}Point estimate} & {\footnotesize{}$-2.35$} & {\footnotesize{}$-2.54$} & {\footnotesize{}$-2.39$} & {\footnotesize{}$-2.22$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
{\footnotesize{}$p/n=0.018$} & {\footnotesize{}$h/b$ unrestricted} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-7.25,2.48]$} & {\footnotesize{}$[-7.25,1.28]$} & {\footnotesize{}$[-6.75,1.6]$} & {\footnotesize{}$[-6.93,2.25]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-12.36$} & {\footnotesize{}$-14.24$} & {\footnotesize{}$-5.74$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.337} & {\footnotesize{}0.170} & {\footnotesize{}0.227} & {\footnotesize{}0.317}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h/b=1$} &  &  &  & \\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust 95\% CI} & {\footnotesize{}$[-10.22,1.42]$} & {\footnotesize{}$[-9.97,-0.03]$} & {\footnotesize{}$[-9.74,0.03]$} & {\footnotesize{}$[-10,1.07]$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}CI length change (\%)} &  & {\footnotesize{}$-14.58$} & {\footnotesize{}$-16.10$} & {\footnotesize{}$-4.94$}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}\hspace{3mm}Robust p-value} & {\footnotesize{}0.139} & {\footnotesize{}0.049} & {\footnotesize{}0.051} & {\footnotesize{}0.110}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$h$, $b$} & {\footnotesize{}6.37, 9.16} & {\footnotesize{}6.37, 9.16} & {\footnotesize{}6.58, 9.88} & {\footnotesize{}4.31, 6.82}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}$n_{-}$, $n_{+}$} & {\footnotesize{}60, 56} & {\footnotesize{}61, 56} & {\footnotesize{}63, 58} & {\footnotesize{}61, 56}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
 & {\footnotesize{}Selected covariates} &  &  &  & {\footnotesize{}\% of adult population}\\}

\usepackage{amsfonts}

\newtheorem{thm}{Theorem}
\newtheorem{lem}{Lemma}[section]
\newtheorem{prop}{Proposition}
\newtheorem{asm}{Assumption}
\theoremstyle{definition}
\newtheorem{rem}{Remark}

\pdfminorversion=4

\makeatother

\begin{document}
\title{Regression Discontinuity Design with Potentially Many Covariates\thanks{This paper is a developed version of the previous manuscript (https://sticerd.lse.ac.uk/dps/em/em601.pdf)
inspired by the discussion with Matias Cattaneo. We are grateful to
Matias Cattaneo and Sebastian Calonico for helpful comments and discussions.
This research was supported by Grants-in-Aid for Scientific Research
20K01598 from the Japan Society for the Promotion of Science (Arai)
and the ERC Consolidator Grant (SNP 615882) (Otsu). Financial support
from the Center for National Competitiveness in the Institute of Economic
Research of Seoul National University and the Ministry of Education
of the Republic of Korea and the National Research Foundation of Korea
(NRF-2018S1A5A2A01033487) is gratefully acknowledged (Seo).}}
\author{Yoichi Arai,\thanks{School of Social Sciences, Waseda University, 1-6-1 Nishiwaseda, Shinjuku-ku,
Tokyo 169-8050, Japan. Email: [email removed]}\ \ Taisuke Otsu\thanks{Department of Economics, London School of Economics, Houghton Street,
London, WC2A 2AE, UK. Email: [email removed]}\ \ and Myung Hwan Seo\thanks{Department of Economics, Seoul National University, 1 Gwankro Gwanakgu,
Seoul, 08826, Korea. Email: [email removed]}}
\maketitle
\begin{abstract}
This paper studies the case of possibly high-dimensional covariates
in the regression discontinuity design (RDD) analysis. In particular,
we propose estimation and inference methods for the RDD models with
covariate selection which perform stably regardless of the number
of covariates. The proposed methods combine a localization approach
using kernel weights with $\ell_{1}$-penalization to handle high-dimensional
covariates. We provide theoretical and numerical results which illustrate
the usefulness of the proposed methods. Theoretically, we present
risk and coverage properties for our point estimation and inference
methods, respectively. Under certain special cases, the proposed estimator
becomes more efficient than the conventional covariate adjusted estimator
at the cost of an additional sparsity condition. Numerically, our
simulation experiments and empirical example show the robust behaviors
of the proposed methods to the number of covariates in terms of bias
and variance for point estimation and coverage probability and interval
length for inference.
\end{abstract}

\section{Introduction}

In causal or treatment effect analysis, discontinuities in regression
functions induced by an assignment variable can provide useful information
to identify certain causal effects. The regression discontinuity design
(RDD) has been widely applied in observational studies to identify
the average treatment effect at the discontinuity point. For the RDD,
the causal parameters of interest are identified by some contrasts
of the left and right limits of the conditional mean functions. See
e.g. Imbens and Lemieux (2008), Cattaneo, Titiunik and Vazquez-Bare
(2020), an edited volume by Cattaneo and Escanciano (2017), and references
therein.

In the growing literature on the RDD analysis, this paper focuses
on the RDDs where covariates are included in the estimation, which
are extensively studied by Calonico, Cattaneo, Farrell and Titiunik
(2019) (hereafter, CCFT). See also Fr�lich and Huber (2019) for an
alternative estimation method based on kernel smoothing after localization
around the cutoff. In practice, researchers often augment the regression
models for the RDD analysis with various additional predetermined
covariates such as demographic or socioeconomic characteristics for
data units. For several RDD estimators using covariates based on local
polynomial regression methods, CCFT investigated the MSE expansion,
asymptotic efficiency, and data-driven bandwidth selection methods.
Furthermore, CCFT developed asymptotic distributional approximations
for those estimators and proposed valid inference procedures by constructing
bias and variance estimators with covariate adjustment. These results
may be considered as extensions of the analyses in Calonico, Cattaneo
and Titiunik (2014) (hereafter, CCT) combined with robust bias correction
methods in Calonico, Cattaneo and Farrell (2018, 2020) to incorporate
covariates in the RDD analysis. See also Calonico, Cattaneo, Farrell
and Titiunik (2017) for a statistical package on these methods.

In randomized controlled trials, regression adjustment using covariates
is a common practice since it is always helpful to improve asymptotic
efficiency of the causal effect estimator as far as a full set of
treatment-covariate interactions is included (Lin, 2013). Also a recent
paper by Lei and Ding (2021) proposed a bias correction method for
the regression adjustment estimator with a diverging number of covariates.
On the other hand, in the RDD analysis, which is a quasi-experiment
setup, the efficiency gain by introducing covariates is not necessarily
guaranteed, and CCFT provided a concrete guideline by clarifying the
conditions to achieve consistency and efficiency gain for the covariate
adjusted RDD estimator. Typically the efficiency improves when the
projection coefficients of the covariates on the outcome are equal
for both control and treatment groups. Since practitioners also commonly
incorporate covariates for the RDD analysis, CCFT's guideline has
a large impact in applied research. When we use covariates, it is
common to employ their transformations and interactions, and the number
of these terms can be pretty large. This paper adds a further guideline
for practitioners who also face a large number of covariates. To begin
with, the (weighted) OLS estimation in CCFT is not applicable when
the number of covariates is larger than the sample size. Also, in
the above scenario for efficiency improvement, it is beneficial to
augment CCFT's procedure with covariate selection by high-dimensional
statistical methods particularly when the regression coefficients
for the conditional mean function satisfy certain sparsity.

For point estimation on the causal effect parameter identified by
the RDD, we consider the Lasso estimator and its post-selection estimator
based on the local linear regression (i.e., eq. (2) of CCFT). The
combination of localization using kernel weights and $\ell_{1}$-penalization
to deal with high-dimensional covariates is particularly relevant
for the RDD analysis, where the effective sample size would be typically
small due to the localization so that the effect of dimensionality
of covariates becomes severer. Theoretically, we derive the $\ell_{1}$-risk
properties of our local Lasso estimator and its post-selection version.
Practically, based on our simulation study, we recommend the CCFT
estimator after selecting covariates by the $\ell_{1}$-penalization
even for a relatively small number of covariates, which exhibits desirable
MSE properties and stability across different setups.

For inference, we propose to select covariates with the local Lasso.
We show that the inference based on the selected covariates can be
implemented in the same manner as in CCFT. We also show that when
the effect of the additional covariates on the potential outcomes
with or without treatment is invariant, our approach can lead to improved
efficiency at the cost of an additional sparsity condition. This sparsity
condition is trivially satisfied when the set of active covariates
is unknown but fixed. Our simulation results demonstrate that our
post-selection confidence interval exhibits robust performances in
terms of both coverages and lengths, even for a relatively small number
of covariates.

This paper also contributes to the large literature on high-dimensional
methods in econometrics and statistics (see, e.g., B�hlmann and van
de Geer, 2011, and Belloni \emph{et al.}, 2018, for an overview) by
combining the kernel localization with $\ell_{1}$-penalization to
handle high-dimensional covariates. Our inference problem can be formulated
as the one for low-dimensional parameters in high-dimensional models.
In statistics literature, many papers investigated this issue, such
as Belloni, Chernozhukov and Hansen (2014), van de Geer, \emph{et
al.} (2014), and Zhang and Zhang (2014). However, these approaches
are not directly applicable to the RDD context because the current
problem concerns the inference on a jump in a nonparametric regression
model.\footnote{A recent paper by Krei\ss\ and Rothe (2023) investigates a similar
estimator to ours, and discusses an inference method based on the
approach by Armstrong and Koles�r (2018).}

This paper is organized as follows. Section \ref{sub:setup} introduces
our basic setup and local Lasso estimator, and presents the $\ell_{1}$-risk
properties. In Section \ref{sub:inf}, we discuss the validity of
CCFT's inference after selecting covariates by our Lasso procedure.
Section \ref{sec:dis} provides discussions on some extensions. A
step-by-step procedure for implementation of our method is described
in Section \ref{sec:rec}. To illustrate the proposed method, Section
\ref{sec:sim} conducts a simulation study, and Section \ref{sec:emp}
presents an empirical example based on the Head Start data.

\section{Main result\label{sec:main}}

\subsection{Setup and local Lasso estimator for covariate selection\label{sub:setup}}

In this subsection, we present our basic setup and introduce the local
Lasso estimator for the RDD with possibly high-dimensional covariates.
For each unit $i=1,\ldots,n$, we observe an indicator variable $T_{i}$
for a treatment ($T_{i}=1$ if treated and $T_{i}=0$ otherwise),
and outcome $Y_{i}=Y_{i}(0)\cdot(1-T_{i})+Y_{i}(1)\cdot T_{i}$, where
$Y_{i}(0)$ and $Y_{i}(1)$ are potential outcomes for $T_{i}=0$
and $T_{i}=1$, respectively. Note that we cannot observe $Y_{i}(0)$
and $Y_{i}(1)$ simultaneously. Our purpose is to make inference on
the causal effect of the treatment, or more specifically, some distributional
aspects of the difference of the potential outcomes $Y_{i}(1)-Y_{i}(0)$.
The RDD analysis focuses on the case where the treatment assignment
$T_{i}$ is completely or partly determined by some observable covariate
$X_{i}$, called the running variable. For example, to study the effect
of class size on pupils' achievements, it is reasonable to consider
the following setup: the unit $i$ is school, $Y_{i}$ is an average
exam score, $T_{i}$ is an indicator variable for the class size ($T_{i}=0$
for one class and $T_{i}=1$ for two classes), and $X_{i}$ is the
number of enrollments. For more examples, see e.g. Imbens and Lemieux
(2008), Cattaneo, Titiunik and Vazquez-Bare (2020), Cattaneo and Escanciano
(2017), and references therein.

Depending on the assignment rule for $T_{i}$ based on $X_{i}$, we
have two cases, called the sharp and fuzzy RDDs. In this section,
we focus on the sharp RDD and discuss the fuzzy RDD in Section \ref{sub:fuzzy}.
In the sharp RDD, the treatment is deterministically assigned as $T_{i}=\mathbb{I}\{X_{i}\geq\bar{x}\}$,
where $\mathbb{I}\{\cdot\}$ is the indicator function and $\bar{x}$
is a known discontinuity (cutoff) point. Throughout the paper, we
normalize $\bar{x}=0$ to simplify the presentation. A parameter of
interest, in this case, is the average causal effect at the discontinuity
point:
\begin{equation}
\tau=\mathbb{E}[Y_{i}(1)-Y_{i}(0)|X_{i}=0].\label{eq:tau}
\end{equation}
Since the difference $Y_{i}(1)-Y_{i}(0)$ is unobservable, we need
a tractable representation of $\tau$ in terms of quantities that
can be estimated by data. If the conditional mean functions $\mathbb{E}[Y_{i}(1)|X_{i}=x]$
and $\mathbb{E}[Y_{i}(0)|X_{i}=x]$ are continuous at the cutoff point
$x=0$, then the average causal effect $\tau$ can be identified as
a contrast of the left and right limits of the conditional mean $\mathbb{E}[Y_{i}|X_{i}=x]$
at $x=0$, that is
\begin{equation}
\tau=\lim_{x\downarrow0}\mathbb{E}[Y_{i}|X_{i}=x]-\lim_{x\uparrow0}\mathbb{E}[Y_{i}|X_{i}=x].\label{eq:tau1}
\end{equation}

As argued in CCFT, it is usually the case that practitioners have
access to additional covariates (denoted by $Z_{i}\in\mathbb{R}^{p}$)
and augment their empirical models with $Z_{i}$ to estimate the causal
effect $\tau$ of interest. This practically relevant setup is extensively
studied in CCFT for the case where $Z_{i}$ is low-dimensional. In
this paper, we consider the case of possibly high-dimensional $Z_{i}$,
and propose a new point estimation method for $\tau$ and an adjustment
of CCFT's inference method.

We examine the case where the additional covariates $Z_{i}$ are predetermined
in the sense that $Z_{i}=Z_{i}(0)\cdot(1-T_{i})+Z_{i}(1)\cdot T_{i}$
but $Z_{i}(0)=_{d}Z_{i}(1)$ for the potential covariates $Z_{i}(0)$
and $Z_{i}(1)$ for $T_{i}=0$ and $T_{i}=1$, respectively. Motivated
by CCFT's recommended model (in their eq. (2)), we propose the local
Lasso estimator $\hat{\theta}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}^{\prime})^{\prime}$
that solves
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},\gamma}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{i}^{\prime}\gamma)^{2}+\lambda_{n}|\gamma|_{1},\label{eq:lasso}
\end{equation}
where $|\gamma|_{1}=\sum_{j=1}^{p}|\gamma_{j}|$ is the $\ell_{1}$-norm
of $\gamma$, $\gamma_{j}$ means the $j$-th element of $\gamma$,
$K(\cdot)$ is a kernel function, $b_{n}$ is a bandwidth, and $\lambda_{n}$
is a penalty level. Popular choices for $K(\cdot)$ are the uniform
and triangular kernels supported on $[-b_{n},b_{n}]$. Based on (\ref{eq:lasso}),
our point estimator for $\tau$ is given by $\hat{\tau}$.

Our preliminary simulation results suggest that the local Lasso estimator
for $\tau$ is somewhat biased in finite samples. Therefore, our recommendation
for point estimation is to employ a post-selection method. Let $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$
for a non-negative sequence $\{\zeta_{n}\}$, and $Z_{\hat{S},i}$
be a subvector of $Z_{i}$ selected by $\hat{S}$. Then the local
post-Lasso estimator $\bar{\theta}=(\bar{\alpha},\bar{\tau},\bar{\beta}_{-},\bar{\beta}_{+},\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$
is defined as a solution of the local least square:
\begin{equation}
\min_{\alpha,\tau,\beta_{-},\beta_{+},t}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)(Y_{i}-\alpha-T_{i}\tau-X_{i}\beta_{-}-T_{i}X_{i}\beta_{+}-Z_{\hat{S},i}^{\prime}t)^{2},\label{eq:post-lasso}
\end{equation}
where $h_{n}$ is another bandwidth, and the estimator for $\tau$
is given by $\bar{\tau}$.

Several points are worthy of remark for this estimator. First, without
the $\ell_{1}$-penalization, our estimator reduces to the local linear-type
estimator recommended by CCFT's eq. (2). Therefore, the proposed estimator
is a natural generalization of CCFT's when the dimension of $Z_{i}$
is high. Second, without the kernel weights for localization, our
estimator in (\ref{eq:lasso}) reduces to the conventional Lasso estimator.
However, since our parameter of interest $\tau$ is identified as
a local object in (\ref{eq:tau1}), it is crucial to introduce such
localization to avoid misspecification bias of the conditional mean
functions. Third, it is often the case that the kernel function $K(\cdot)$
has bounded support. In this case, the effective sample size would
be typically of orders $nb_{n}$ and $nh_{n}$. Thus even if the dimension
of $Z_{i}$ is relatively small compared to the original sample size
$n$, the $\ell_{1}$-penalization would be useful especially for
small values of $b_{n}$ and $h_{n}$. Finally, the trimming term
$\zeta_{n}$ to obtain the set $\hat{S}$ is introduced to stabilize
numerical results (see (\ref{eq:zeta}) below for our recommended
choice based on simulation studies), and theoretically we may set
as $\zeta_{n}=0$.

We now present risk properties of the local Lasso estimators $\hat{\theta}$
and $\bar{\theta}$. Let $G_{i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{i}^{\prime})^{\prime}$
be the vector of regressors in (\ref{eq:lasso}), $G_{i,j}$ be the
$j$-th element of $G_{i}$, and $\Theta_{n}=\arg\min_{\theta}\mathbb{E}[K(X_{i}/b_{n})(Y_{i}-G_{i}^{\prime}\theta)^{2}]$
be an argmin set. We impose the following assumptions.

\begin{asm}\label{asm:point-est} There exists a sequence $\theta_{n}^{*}\in\Theta_{n}$
that satisfies the following conditions.
\begin{enumerate}
\item Let $\epsilon_{i}=\sqrt{K(X_{i}/b_{n})}(Y_{i}-G_{i}^{\prime}\theta_{n}^{*})$.
There exists some $C\in(0,\infty)$ such that
\[
\mathbb{E}[|K(X_{i}/b_{n})G_{i,j}\epsilon_{i}|^{m}]\leq b_{n}m!C^{m-2}/2,
\]
for all $j=1,\ldots,p$ and $m=2,3,\ldots$.
\item Let $\delta_{A}$ be the subvector of $\delta$ for an index set $A$,
$S^{*}=\{j:\theta_{n,j}^{*}\neq0\}$, and $(S^{*})^{c}=\{j:\theta_{n,j}^{*}=0\}$,
where $\theta_{n,j}^{*}$ is the $j$-th element of $\theta_{n}^{*}$.
There exists some $\phi^{*}\in(0,\infty)$ such that
\[
\frac{s^{*}}{|\delta_{S^{*}}|_{1}^{2}}\cdot\min_{\delta:|\delta_{(S^{*})^{c}}|_{1}\leq3|\delta_{S^{*}}|_{1}}\delta^{\prime}\left(\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)G_{i}G_{i}^{\prime}\right)\delta\geq(\phi^{*})^{2},
\]
with probability approaching one, where $s^{*}=|S^{*}|$.
\end{enumerate}
\end{asm}

\begin{asm} \label{asm:kb} $K:\mathbb{R\to\mathbb{R}}$ is a bounded
and symmetric second-order kernel function which is continuous with
a compact support. The bandwidth $b_{n}$ is a positive sequence satisfying
$b_{n}\to0$ and $nb_{n}\to\infty$ as $n\to\infty$. \end{asm}

Assumption \ref{asm:point-est} defines $\theta_{n}^{*}$ as an approximate
linear predictor or the linear projection on the set of included variables
in the index set $S^{*}$ since $\mathbb{E}[K(X_{i}/b_{n})G_{i,j}\epsilon_{i}]=0$
for all $j\in S^{*}$. This assumption is general enough to cover
the setup in CCFT, which assumes $p$ is fixed. In RDD analyses, it
is common to introduce many generated covariates, such as transformations
of initial covariates like polynomials, interactions, and various
basis functions, without knowing which of them are relevant a priori.\footnote{To motivate the use of generated covariates, it is insightful to note
that the asymptotic variance of CCFT's RDD estimator is proportional
to $\mathrm{Var}(\{(Y_{i}(1)-Y_{i}(0))-(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma\}^{2}|X_{i}=0)$
for some $\gamma$, which is considered as the (conditional) variance
of the linear projection error. Although CCFT considered linear projection
due to the constraint on dimensionality, it is clear that the asymptotic
variance is minimized by employing the conditional expectation $\mathbb{E}[Y_{i}(1)-Y_{i}(0)|Z_{i}(1)-Z_{i}(0),X_{i}=0]$
instead of the linear projection $(Z_{i}(1)-Z_{i}(0))^{\prime}\gamma$.
Therefore, it is natural to extend CCFT's approach to high-dimensional
settings by employing generated covariates or series approximation
for the conditional mean.}

Although it is beyond the scope of this paper, the definition of $\theta_{n}^{*}$
could be modified to be an approximate minimizer which does not belong
to $\Theta_{n}$ but gets closer to it at some suitable rate. Since
it complicates the exposition and derivation as in Krei\ss\ and
Rothe (2023) or Belloni, Chernozhukov and Hansen (2014), we maintain
this exact sparsity assumption. For example, such an extension for
approximate sparsity will be useful to allow the situation where the
conditional mean satisfies \emph{$E[Y|X,Z]=E[Y|X,Z_{\mathcal{S}}]$}
for some sparse set $\mathcal{S}$ but the conditional mean function
$E[Y|X,Z_{\mathcal{S}}]$ is nonlinear in $Z_{\mathcal{S}}$ so that
the exact sparsity assumption typically fails.

Assumption \ref{asm:point-est} (1) contains a set of moment conditions,
which are introduced to verify a local-type Bernstein inequality in
Lemma \ref{lem:Bernstein}. The extra factor $b_{n}$ is due to the
presence of the kernel weight $K(X_{i}/b_{n})$. Assumption \ref{asm:point-est}
(2) is a localized version of the compatibility condition. A sufficient
condition for this is the so-called restricted eigenvalue condition.
More specifically, $\min_{\beta:|\beta|_{0}\leq s^{*}}\frac{1}{nb_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{b_{n}}\right)\frac{\beta^{\prime}G_{i}G_{i}^{\prime}\beta}{\beta^{\prime}\beta}$,
where $|\beta|_{0}$ denotes the cardinality of $\beta$, provides
a lower bound for the compatibility constant $(\phi^{*})^{2}$ (see,
e.g., Section 6.13 of B�hlmann and van der Geer, 2011). If CCFT's
method is feasible with each subset of covariates of dimension $2s^{*}$,
then the restricted eigenvalue condition is indeed satisfied. Assumption
\ref{asm:kb} contains assumptions on the kernel $K$ and bandwidth
$b_{n}$, which are standard in the literature of nonparametric methods.
Note that since $\theta_{n}^{*}$ and $S^{*}$ depend on $b_{n}$,
Assumption \ref{asm:point-est} should be satisfied along each sequence
$\{b_{n}\}$. Also, the deviation bounds on the prediction and estimation
errors of $\hat{\theta}$ will be given as functions of $s^{*}$.
While a precise condition on $s^{*}$ is hard to specify and depends
on the sampling distribution, it would be typically smaller order
than $\sqrt{nb_{n}}$ to satisfy the compatibility condition in Assumption
\ref{asm:point-est} (2).

Let $\hat{\gamma}_{\hat{S}}$ be the subvector of $\hat{\gamma}$
selected by $\hat{S}$, $\hat{\theta}_{\hat{S}}=(\hat{\alpha},\hat{\tau},\hat{\beta}_{-},\hat{\beta}_{+},\hat{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$Z_{\hat{S},i}$ be the subvector of $Z_{i}$ selected by $\hat{S}$,
$G_{\hat{S},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S},i}^{\prime})^{\prime}$,
and $m_{n}=\lambda_{\min}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}\right)^{-1}$,
where $\lambda_{\min}(A)$ means the minimum eigenvalue of a matrix
$A$. Also let $\hat{S}_{1}=\{i:0<|\hat{\gamma}_{i}|<\zeta_{n}\}$,
$S_{n}=\hat{S}\cup\hat{S}_{1}=\{i:\hat{\gamma}_{i}\neq0\}$, $Z_{\hat{S}_{1},i}$
be the subvector of $Z_{i}$ selected by $\hat{S}_{1}$, $G_{\hat{S}_{1},i}=(1,T_{i},X_{i},T_{i}X_{i},Z_{\hat{S}_{1},i}^{\prime})^{\prime}$,
and $m_{1n}=\lambda_{\max}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/h_{n})G_{\hat{S}_{1},i}G_{\hat{S}_{1},i}^{\prime}\right)$,
where $\lambda_{\max}(A)$ means the maximum eigenvalue of a matrix
$A$. The $\ell_{1}$-risk properties of the local Lasso and post-Lasso
estimators (for the case of $h_{n}=b_{n}$) are obtained as follows.

\begin{thm}\label{thm:point-est} Suppose $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$.
\begin{description}
\item [{(i)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb}, it
holds
\begin{equation}
|\hat{\theta}-\theta_{n}^{\ast}|_{1}\leq C\frac{\lambda_{n}s^{*}}{\phi^{*2}},\label{eq:l1b}
\end{equation}
for some $C\in(0,\infty)$ with probability approaching one.
\item [{(ii)}] Under Assumptions \ref{asm:point-est}-\ref{asm:kb} and
$h_{n}=b_{n}$, it holds
\begin{equation}
|\bar{\theta}_{S_{n}}-\hat{\theta}_{S_{n}}|_{1}\leq(m_{n}|S_{n}|\lambda_{n})\vee(|\hat{S}_{1}|\zeta_{n})\vee(m_{1n}|\hat{S}_{1}|\zeta_{n}^{2}/\lambda_{n}).\label{eq:l1c}
\end{equation}
\end{description}
\end{thm}

The proof of this theorem is presented in Appendix \ref{sub:pf1}.
This theorem characterizes the risk properties of the estimators $\hat{\theta}$
and $\bar{\theta}$ around $\theta_{n}^{*}$. The risk bound of $\hat{\theta}$
depends on the tuning parameter $\lambda_{n}$, the number of non-zero
coefficients $s^{*}$, and the compatibility constant $\phi^{*}$.
Note that the decay rate of $\lambda_{n}$ is bounded from below by
$\sqrt{\log p/(nb_{n})}$. Thus, the risk bound of $\hat{\theta}$
gets worse as the number of covariates $p$ increases or the effective
sample size $nb_{n}$ due to the kernel localization decreases. The
result (\ref{eq:l1c}) for the post-selection estimator $\bar{\theta}$
shows that the deviation from the original Lasso estimator $\hat{\theta}$
is small when tuning parameter $\lambda_{n}$ or the number of selected
covariates $|S_{n}|$ is small, or the minimum eigenvalue of $\frac{1}{nh_{n}}\sum_{i=1}^{n}K(X_{i}/b_{n})G_{\hat{S},i}G_{\hat{S},i}^{\prime}$
is large. If the trimming parameter $\zeta_{n}$ is of similar magnitude
of $\lambda_{n}$, then the three terms in the bounds are of similar
magnitude. If $\zeta_{n}$ is smaller order of magnitude than $\lambda_{n}$,
then the first term will dominate. We suggest some practical choice
of the trimming term $\zeta_{n}$ in Section \ref{sec:sim} based
on our simulation studies.

The above theorem is on estimation of the coefficients of the best
linear predictor $\theta_{n}^{*}$ defined in Assumption \ref{asm:point-est}
(1). Additionally suppose that the assumptions of Lemma 1 of CCFT
hold true, and the covariates $Z_{i}$ are predetermined. Then we
can guarantee that the second element of $\theta_{n}^{*}$ coincides
with the average causal effect $\tau$ in (\ref{eq:tau}) so that
Theorem \ref{thm:point-est} provides the conditions for the consistency
and convergence rate of $\hat{\tau}$ to $\tau$. If $Z_{i}$ are
not predetermined (i.e., $Z_{i}(0)\neq_{d}Z_{i}(1)$), then $\hat{\tau}$
typically converges to $\tau$ minus some bias component, which is
obtained as a limit of CCFT's bias term in their Lemma 1.

Our estimators and above theorem can be extended to other regression
models that contain the covariates $\{T_{i}Z_{i},(1-T_{i})Z_{i}\}$,
$(Z_{i}-\bar{Z})$, or $\{T_{i}(Z_{i}-\bar{Z}),(1-T_{i})(Z_{i}-\bar{Z})\}$
as in CCFT. However, as shown in Lemma 1 of CCFT, such estimators
require more stringent conditions to guarantee the consistency for
$\tau$. Furthermore, the local Lasso regression (\ref{eq:lasso})
can be extended to incorporate polynomials of $X_{i}$ and $T_{i}X_{i}$
even though this paper focuses on the local linear model.

Finally, we discuss the choices of the localization bandwidths $b_{n}$
and $h_{n}$ and regularization parameter $\lambda_{n}$. We can use
the MSE-optimal bandwidth based on the suggestion by CCFT and the
regularization parameter $\lambda_{n}$ using cross-validation by
Friedman, Hastie and Tibshirani (2010) or a data-driven choice by
Belloni, Chernozhukov and Hansen (2014) among others. See Theorem
\ref{thm:t} in the next subsection for their justification, and Section
\ref{sec:rec} for a detail on our practical recommendation.

\subsection{Inference\label{sub:inf}}

We next consider interval estimation and hypothesis testing on the
average causal effect $\tau$. For finite or low-dimensional $Z_{i}$,
we recommend to use CCFT's bias corrected inference method. This subsection
argues that we can still apply CCFT's inference procedure for high-dimensional
$Z_{i}$, provided that CCFT's conditions remain valid for $S^{*}$
and the subvector $\theta_{n,S^{*}}^{*}$ of $\theta_{n}^{*}$ selected
by $S^{*}$.

In this subsection, we specify the tuning constant $\zeta_{n}$ to
obtain $\hat{S}=\{j:|\hat{\gamma}_{j}|\ge\zeta_{n}\}$ as
\begin{equation}
\zeta_{n}=\lambda_{n}\varrho_{n}\sum_{j=1}^{p}\mathbb{I}\{\hat{\gamma}_{j}\neq0\},\label{eq:zeta}
\end{equation}
where we set $\varrho_{n}=\log\log\log n$. This choice of $\varrho_{n}$
is based on the simulation experiments in Section \ref{sec:sim},
and it is not shown to be optimal but works reasonably well. Based
on the selected covariates by $\hat{S}$ with $\zeta_{n}$ in (\ref{eq:zeta}),
we apply CCFT's bias corrected t-ratio to conduct statistical inference
on the causal effect parameter $\tau$.

Consider the local post-Lasso estimator $\bar{\tau}$ defined by (\ref{eq:post-lasso}).
As shown in Appendix \ref{sub:pf2}, the dominant term of $\bar{\tau}$
can be characterized as
\begin{equation}
\bar{\tau}=e_{2}^{\prime}\left(\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}G_{1i}^{\prime}\right)^{-1}\frac{1}{nh_{n}}\sum_{i=1}^{n}K\left(\frac{X_{i}}{h_{n}}\right)G_{1i}\xi_{i}+o_{p}((nh_{n})^{-1/2}),\label{eq:lin}
\end{equation}
where $e_{2}=(0,1,0,0)^{\prime}$, $G_{1i}=(1,T_{i},X_{i},T_{i}X_{i})^{\prime}$,
$\xi_{i}=T_{i}\xi_{i}(1)+(1-T_{i})\xi_{i}(0)$ with $\xi_{i}(t)=Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}$,
and
\begin{equation}
\gamma_{Y}=\arg\min_{\gamma:\gamma_{j}=0\text{ for }j\notin S^{*}}\mathbb{E}[(\tilde{Y}-Z^{\prime}\gamma)^{2}|X=0],\label{eq:gamY}
\end{equation}
with $\tilde{Y}=Y(1)-Y(0)-\mathbb{E}[Y(1)-Y(0)|X=0]$. Indeed, the
asymptotic linear form in (\ref{eq:lin}) is analogous to the one
derived for the case of fixed dimensional $Z_{i}$ in CCFT (except
that $\gamma_{Y,j}=0$ for $j\notin S^{*}$). Therefore, the pre-asymptotic
bias and variance of $\bar{\tau}$ can be analogously written as
\begin{eqnarray}
\mathcal{B} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(q^{\prime}\mu_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(q^{\prime}\mu_{+}^{(2)}),\nonumber \\
\mathcal{V} & = & (q\otimes P_{-}^{\prime}e_{1})^{\prime}\Sigma_{-}(q\otimes P_{-}^{\prime}e_{1})+(q\otimes P_{+}^{\prime}e_{1})^{\prime}\Sigma_{+}(q\otimes P_{+}^{\prime}e_{1}),\label{eq:V}
\end{eqnarray}
respectively, where $e_{1}=(1,0)^{\prime}$, $R=\left[\begin{array}{ccc}
1 & \ldots & 1\\
X_{1}/h_{n} & \ldots & X_{n}/h_{n}
\end{array}\right]^{\prime}$, $q=(1,-\gamma^{*\prime})^{\prime}$, and
\begin{eqnarray}
K_{-} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}<0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}<0\}K(X_{n}/h_{n})),\nonumber \\
K_{+} & = & h_{n}^{-1}\mathrm{diag}(\mathbb{I}\{X_{1}\ge0\}K(X_{1}/h_{n}),\ldots,\mathbb{I}\{X_{n}\ge0\}K(X_{n}/h_{n})),\nonumber \\
\Gamma_{-} & = & n^{-1}R^{\prime}K_{-}R,\qquad\Gamma_{+}=n^{-1}R^{\prime}K_{+}R,\nonumber \\
\vartheta_{-} & = & n^{-1}R^{\prime}K_{-}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\qquad\vartheta_{+}=n^{-1}R^{\prime}K_{+}[X_{1}^{2}/h_{n}^{2},\ldots,X_{n}^{2}/h_{n}^{2}]^{\prime},\nonumber \\
\mu_{-}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(0)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
\mu_{+}^{(2)} & = & \left[\left.\frac{\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]}{\partial x^{2}}\right|_{x=0},\left.\frac{\partial^{2}\mathbb{E}[Z_{i}^{*}(1)^{\prime}|X_{i}=x]}{\partial x^{2}}\right|_{x=0}\right]^{\prime},\nonumber \\
P_{-} & = & \sqrt{\frac{h}{n}}\Gamma_{-}^{-1}R^{\prime}K_{-},\qquad P_{+}=\sqrt{\frac{h}{n}}\Gamma_{+}^{-1}R^{\prime}K_{+},\nonumber \\
\Sigma_{-} & = & \mathrm{Var}(\mathrm{vec}(\mathbf{Y}(0),\mathbf{Z}^{*}(0))|\mathbf{X}=0),\qquad\Sigma_{+}=\mathrm{Var}(\mathrm{vec}(\mathbf{Y}(1),\mathbf{Z}^{*}(1))|\mathbf{X}=0),\label{eq:Vnote}
\end{eqnarray}
with $\mathbf{Y}(0)=(Y_{1}(0),\ldots,Y_{n}(0))^{\prime}$, $\mathbf{Y}(1)=(Y_{1}(1),\ldots,Y_{n}(1))^{\prime}$,
$\mathbf{X}=(X_{1},\ldots,X_{n})^{\prime}$, $\mathbf{Z}^{*}(0)=(Z_{S^{*},1}(0),\ldots,Z_{S^{*},n}(0))^{\prime}$,
and $\mathbf{Z}^{*}(1)=(Z_{S^{*},1}(1),\ldots,Z_{S^{*},n}(1))^{\prime}$.

By estimating the unknown components, the pre-asymptotic bias and
variance can be estimated as
\begin{eqnarray*}
\bar{\mathcal{B}} & = & \frac{1}{2}e_{1}^{\prime}\Gamma_{-}^{-1}\vartheta_{-}(\bar{q}^{\prime}\bar{\mu}_{-}^{(2)})+\frac{1}{2}e_{1}^{\prime}\Gamma_{+}^{-1}\vartheta_{+}(\bar{q}^{\prime}\bar{\mu}_{+}^{(2)}),\\
\bar{\mathcal{V}} & = & (\bar{q}\otimes P_{-}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{-}(\bar{q}\otimes P_{-}^{\prime}e_{1})+(\bar{q}\otimes P_{+}^{\prime}e_{1})^{\prime}\bar{\Sigma}_{+}(\bar{q}\otimes P_{+}^{\prime}e_{1}),
\end{eqnarray*}
respectively, where $\bar{q}^{\prime}=(1,-\bar{\gamma}_{\hat{S}}^{\prime})^{\prime}$,
$\bar{\mu}_{-}^{(2)}$ and $\bar{\mu}_{+}^{(2)}$ are local polynomial
estimators of $\mu_{-}^{(2)}$ and $\mu_{+}^{(2)}$ for the elements
corresponding to $Z_{\hat{S},i}$, respectively, and $\bar{\Sigma}_{+}$
and $\bar{\Sigma}_{-}$ are conditional variance estimators of $\Sigma_{-}$
and $\Sigma_{+}$, respectively, such as the nearest neighborhood
or plug-in estimator in Section 7.9 of CCFT's supplement. Based on
these estimators, the t-ratio for $\tau$ is obtained as
\begin{equation}
T_{\tau}=\frac{\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}-\tau}{\sqrt{(nh_{n})^{-1}\bar{\mathcal{V}}}},\label{eq:t}
\end{equation}
which is exactly the same as the t-ratio in Theorem 2 of CCFT but
using the selected covariates $Z_{\hat{S},i}$. By extending the theoretical
developments in CCFT, we obtain the following result.

\begin{thm} \label{thm:t} Suppose Assumptions \ref{asm:point-est}-\ref{asm:kb}
hold true. Suppose for all $x$ in a neighborhood of $0$ and $t=0,1$,
the density of $X_{i}$ is continuous and bounded away from zero,
$\mathbb{E}[(Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x]$ is three times
continuously differentiable, $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
$\mathbb{E}[Z_{i}(t)Y_{i}(t)|X_{i}=x]$ is continuously differentiable,
$\mathrm{Var}((Y_{i}(t),Z_{i}(t)^{\prime})|X_{i}=x)$ is continuously
differentiable and invertible, and $\mathbb{E}[|(Y_{i}(t),Z_{i}(t)^{\prime})|^{4}|X_{i}=x]$
is continuous. For all $j,k$ and positive integers $m$, and some
finite $C$, $\mathbb{E}[|K(X_{i}/h_{n})G_{1i,j}Z_{i,k}|^{m}]\leq h_{n}m!C^{m-2}/2$.
Finally, assume $h_{n}\to0$, $nh_{n}\to\infty$, $\lambda_{n}^{-1}\sqrt{\log p/(nb_{n})}\to0$,
and $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$.
Then the conditional MSE expansion of the first term in (\ref{eq:lin})
(denoted by $\bar{\tau}_{1}$) is obtained as
\begin{equation}
\mathbb{E}[(\bar{\tau}_{1}-\tau)^{2}|\mathbf{X}]=h_{n}^{4}\mathcal{B}^{2}\{1+o_{p}(1)\}+\frac{1}{nh_{n}}\mathcal{V}.\label{eq:MSE}
\end{equation}
Furthermore, if we additionally assume $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$, then

\begin{eqnarray}
T_{\tau} & \overset{d}{\to} & N(0,1).\label{eq:tau-t}
\end{eqnarray}
 \end{thm}

The results in (\ref{eq:MSE}) and (\ref{eq:tau-t}) are analogous
to CCFT's Theorems 1 and 2, respectively. This theorem theoretically
supports to employ the bias correction and bandwidth selection methods
by CCFT based on the selected covariates $Z_{\hat{S},i}$. See Section
\ref{sec:rec} below for our practical recommendation. The assumption
$\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$
is natural for predetermined covariates but may be relaxed by introducing
additional regularity conditions (see, Krei\ss\ and Rothe, 2023).
Other assumptions except for the last one are also imposed in CCFT.
The assumption $(\sqrt{\log p}+\sqrt{nh_{n}}h_{n}^{2})(h_{n}^{2}+b_{n}^{2}+\lambda_{n}+\zeta_{n}+\zeta_{n}^{2}/\lambda_{n})s^{*}\to0$
is used to control the remainder term in (\ref{eq:lin}), and can
be considered as a sparsity assumption to restrict the growth rate
of $s^{*}$.\footnote{In the standard Lasso literature, we typically impose $\frac{s^{*}\log p}{\sqrt{n}}\to0$
and the minimal penalty level requirement on $\lambda_{n}$. For comparison,
consider the following standard setting for tuning parameters, where
$b_{n}\sim h_{n}\sim n^{-1/5}$, $\lambda_{n}=a_{n}\sqrt{\log p/(nb_{n})}$
with a slowly diverging $a_{n}$ and $\zeta_{n}=O(\lambda_{n})$.
Then the condition on $s^{*}$ reduces to $\frac{s^{*}a_{n}\log p}{\sqrt{nh_{n}}}\to0$,
which is analogous to the standard case.} Although the conditions $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
and $\frac{\bar{\mathcal{V}}}{\mathcal{V}}\overset{p}{\to}1$ are
high level, these are typically satisfied for the bias and variance
estimators discussed in CCFT.\footnote{Under the assumption $\partial^{2}\mathbb{E}[Z_{i}(0)|X_{i}=x]/\partial x^{2}=\partial^{2}\mathbb{E}[Z_{i}(1)|X_{i}=x]/\partial x^{2}$,
the components $q^{\prime}\mu_{-}^{(2)}$ and $q^{\prime}\mu_{+}^{(2)}$
in $\mathcal{B}$ become $\mu_{Y-}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(0)|X_{i}=x]/\partial x^{2}\right|_{x=0}$
and $\mu_{Y+}^{(2)}=\left.\partial^{2}\mathbb{E}[Y_{i}(1)|X_{i}=x]/\partial x^{2}\right|_{x=0}$,
respectively. Thus, in this case, the convergence rates of the conventional
local polynomial estimators for $\mu_{Y-}^{(2)}$ and $\mu_{Y+}^{(2)}$
guarantee $\sqrt{\frac{nh_{n}^{5}}{\mathcal{V}}}(\bar{\mathcal{B}}-\mathcal{B})\overset{p}{\to}0$
(see, e.g., Fan and Gijbels, 1992, and Ruppert and Wand, 1994).} See Remark \ref{rem:V} below for a specific example of the variance
estimator $\bar{\mathcal{V}}$.

\begin{rem} {[}Efficiency comparison{]} It should be noted that the
asymptotic variance $\mathcal{V}$ in (\ref{eq:V}) of the estimator
$\bar{\tau}$ (or the bias corrected version $\bar{\tau}-h_{n}^{2}\bar{\mathcal{B}}$)
takes the same form as the one in CCFT even when the dimension of
$Z_{S^{*}}$ grows as $n$ increases. As investigated in CCFT, we
can see that the relative efficiency of $\bar{\tau}$ compared to
the conventional RDD estimator without covariates (say, $\hat{\tau}_{\mathrm{unadjusted}}$)
is
\begin{equation}
\frac{\mathcal{V}}{\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}}=\frac{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)-Z_{i}(t)^{\prime}\gamma_{Y}|X_{i}=0)}{\sum_{t=0}^{1}\mathrm{Var}(Y_{i}(t)|X_{i}=0)},\label{eq:rel}
\end{equation}
where $\mathcal{\mathcal{V}}_{\mathrm{unadjusted}}$ is the asymptotic
variance of $\hat{\tau}_{\mathrm{unadjusted}}$ and $\gamma_{Y}$
is defined in (\ref{eq:gamY}). Generally there is no clear ranking
for these asymptotic variances. However, letting $\gamma_{Y,S^{*}}$
be an $s^{*}$-dimensional subvector of $\gamma_{Y}$ selected by
$S^{*}$, in an important special case where
\begin{equation}
\gamma_{Y,S^{*}}=\mathrm{Var}(Z_{S^{*},i}(t)|X_{i}=0)^{-1}\mathbb{E}[(Z_{S^{*},i}(t)-\mathbb{E}[Z_{S^{*},i}(t)|X_{i}])Y_{i}(t)|X_{i}=0]\quad\text{for }t=0\text{ and }1,\label{eq:gam}
\end{equation}
the coefficient vector $\gamma_{Y,S^{*}}$ becomes the best linear
approximation by $Z_{S^{*}}(t)$ for each group and thus $\bar{\tau}$
is asymptotically more efficient than $\hat{\tau}_{\mathrm{unadjusted}}$.

The relative efficiency in (\ref{eq:rel}) is also insightful to illustrate
the merit of our Lasso approach compared to CCFT. Let $Z=[Z_{S^{*}}:Z_{S^{*c}}]$
and $Z_{C}$ be covariates employed to apply the CCFT estimator $\hat{\tau}_{\mathrm{CCFT}}$
(without the Lasso covariates selection). Consider the special case
in (\ref{eq:gam}). As far as $Z_{C}$ contains $Z_{S^{*}}$, $\bar{\tau}$
and $\hat{\tau}_{\mathrm{CCFT}}$ achieve the same asymptotic efficiency
$\mathcal{V}$. On the other hand, if $Z_{C}$ does not contain some
elements of $Z_{S^{*}}$, then under (\ref{eq:gam}), the CCFT estimator
$\hat{\tau}_{\mathrm{CCFT}}$ is asymptotically less efficient than
the post-Lasso estimator $\bar{\tau}$. Therefore, when researchers
are less certain whether all elements of $Z_{S^{*}}$ are included
in $Z_{C}$ typically due to too many candidates in $Z$ or too small
effective sample sizes used for estimation, our Lasso-based approach
may be more attractive to achieve asymptotic efficiency $\mathcal{V}$
in more broader situations. We emphasize that such an efficiency gain
of our estimator $\bar{\tau}$ can be achieved at the cost of the
additional sparsity condition in Assumption \ref{asm:point-est},
which is trivially satisfied when the set of active covariates is
unknown but fixed. \end{rem}

\begin{rem} {[}Finite sample comparison{]} Although there is no gain
of using $\bar{\tau}$ instead of $\hat{\tau}_{\mathrm{CCFT}}$ in
terms of asymptotic efficiency as far as $Z_{C}$ contains $Z_{S^{*}}$,
the estimator $\bar{\tau}$ may exhibit better finite sample performance
even in such a scenario. To see this point, let $(\hat{\beta}_{C}^{\prime},\hat{\gamma}_{C}^{\prime})$
be the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{C}$ so that $\hat{\tau}_{\mathrm{CCFT}}$ is the second
element of $\hat{\beta}_{C}$. As can be seen from CCFT's supplement,
one of the remainder terms of $\sqrt{\frac{nh_{n}}{\mathcal{V}}}(\hat{\tau}_{\mathrm{CCFT}}-\tau)$
involves a linear combination of the estimation error $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$.
Since the $\ell_{2}$-convergence rate of $\hat{\gamma}_{C}-(\gamma^{*\prime},0^{\prime})^{\prime}$
is typically of order $\sqrt{\dim Z_{C}/n}$, this remainder term
is of larger order than $\hat{\gamma}^{*}-\gamma_{Y}$, where $(\hat{\beta}^{*\prime},\hat{\gamma}^{*\prime})$
is the OLS estimator for the regression of $K^{1/2}Y$ on $K^{1/2}(1,T,X,TX)$
and $K^{1/2}Z_{S^{*}}$. Even though these remainder terms do not
appear in the first order asymptotic distribution, they contribute
to the finite sample behaviors of $\bar{\tau}$ and $\hat{\tau}_{\mathrm{CCFT}}$.
Another finite sample issue we encounter in our simulation study below
is that as the dimension of $Z_{C}$ increases, the values of the
MSE-optimal bandwidth tend to be smaller (due to larger bias estimates
but relatively stable variance estimates). Thus the effective sample
size used for the RDD estimation tends to be smaller so that we observe
larger variations in the resulting RDD estimates across simulation
draws. Finally, our simulation results in Section \ref{sec:sim} (particularly
DGPs 2 and 3 with large $p$) illustrate that the covariate selection
approach exhibits smaller standard deviations than CCFT in finite
samples even when the asymptotic variances may be equivalent. \end{rem}

\begin{rem} {[}Selection consistency{]} Although Theorem \ref{thm:t}
is our main result on inference of the causal effect $\tau$, it is
also possible to derive the consistency of the selection procedure
(i.e., $\mathbb{P}\{\hat{S}=S^{*}\}\to1$) under some additional $\beta$-min
type condition (i.e., there exists some $\varepsilon>0$ such that
$|\gamma_{j}^{*}|>\lambda_{n}\varrho_{n}s^{*}(1+\varepsilon)$ for
each $j\in S^{*}$). See a working paper version of this paper (Arai,
Otsu and Seo, 2021) for more details on the selection consistency.
\end{rem}

\begin{rem}\label{rem:V} {[}Estimation of $\mathcal{V}${]} An example
of the estimator $\bar{\mathcal{V}}$ for the asymptotic variance
$\mathcal{V}$ is the nearest neighborhood estimator, which is employed
in CCT, CCFT, and our numerical illustrations. Let
\begin{eqnarray*}
\bar{\varepsilon}_{V-,i} & = & \mathbb{I}\{X_{i}<0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{-,j}(i)}\right),\\
\bar{\varepsilon}_{V+,i} & = & \mathbb{I}\{X_{i}\ge0\}\sqrt{\frac{J}{J+1}}\left(V_{i}-\frac{1}{J}\sum_{j=1}^{J}V_{\ell_{+,j}(i)}\right),
\end{eqnarray*}
for $V\in\{Y,Z_{1},\ldots,Z_{p}\}$ and a fixed positive integer $J$,
where $\ell_{-,j}(i)$ is the index of the $j$-th closest unit to
unit $i$ among $\{i:X_{i}<0\}$, and $\ell_{+,j}(i)$ is the index
of the $j$-th closest unit to unit $i$ among $\{i:X_{i}\ge0\}$.
The nearest neighborhood estimators of $\Sigma_{-}$ and $\Sigma_{+}$
in (\ref{eq:Vnote}) are defined as
\[
\bar{\Sigma}_{-}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY-}^{NN} & \bar{\Sigma}_{YZ_{1}-}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}-}^{NN}\\
\bar{\Sigma}_{Z_{1}Y-}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}-}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y-}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}-}^{NN}
\end{array}\right],\quad\bar{\Sigma}_{+}^{NN}=\left[\begin{array}{cccc}
\bar{\Sigma}_{YY+}^{NN} & \bar{\Sigma}_{YZ_{1}+}^{NN} & \cdots & \bar{\Sigma}_{YZ_{|\hat{S}|}+}^{NN}\\
\bar{\Sigma}_{Z_{1}Y+}^{NN} & \bar{\Sigma}_{Z_{1}Z_{1}+}^{NN}\\
\vdots &  & \ddots\\
\bar{\Sigma}_{Z_{|\hat{S}|}Y+}^{NN} &  &  & \bar{\Sigma}_{Z_{|\hat{S}|}Z_{|\hat{S}|}+}^{NN}
\end{array}\right],
\]
where $\bar{\Sigma}_{VW-}^{NN}$ and $\bar{\Sigma}_{VW+}^{NN}$ are
$n\times n$ matrices whose $(i,j)$-th elements are
\begin{eqnarray*}
[\bar{\Sigma}_{VW-}^{NN}]_{i,j} & = & \mathbb{I}\{X_{i}<0\}\mathbb{I}\{X_{j}<0\}\mathbb{I}\{i=j\}\bar{\varepsilon}_{V-,i}\bar{\varepsilon}_{W-,j},\\{
\hline
\end{tabular}{\footnotesize\par}
\end{table}
{\footnotesize\par}

\newpage{}

\begin{figure}[H]
	\caption{Comparison of three approaches ($n=1000$, $45$ covariates) for point
		estimates (left panel) and CI lengths (right panel) }

	\begin{centering}
		\includegraphics[scale=0.4]{fig1a_headstart_pe_change_1000}\includegraphics[scale=0.4]{fig1b_headstart_ci_change_1000}
		\par\end{centering}
	\centering{}{\footnotesize{}\label{fig:1}}{\footnotesize\par}
\end{figure}

\begin{figure}[H]
	\caption{Comparison of three approaches ($n=500$, 45 covariates) for point
		estimates (left panel) and CI lengths (right panel) }

	\begin{centering}
		\includegraphics[scale=0.4]{fig2a_headstart_pe_change_500}\includegraphics[scale=0.4]{fig2b_headstart_ci_change_500}
		\par\end{centering}
	\centering{}{\footnotesize{}\label{fig:2}}{\footnotesize\par}
\end{figure}

\begin{figure}
	\caption{Comparison of three approaches ($n=1000$, 9 covariates) for point
		estimates (left panel) and CI lengths (right panel)}

	\begin{centering}
		\includegraphics[scale=0.4]{fig3a_headstart_pe_change_1000_9}\includegraphics[scale=0.4]{fig3b_headstart_ci_change_1000_9}
		\par\end{centering}
	\centering{}{\footnotesize{}\label{fig:3}}{\footnotesize\par}
\end{figure}

\begin{figure}
	\caption{Comparison of three approaches ($n=500$, 9 covariates) for point
		estimates (left panel) and CI lengths (right panel)}

	\begin{centering}
		\includegraphics[scale=0.4]{fig4a_headstart_pe_change_500_9}\includegraphics[scale=0.4]{fig4b_headstart_ci_change_500_9}
		\par\end{centering}
	\centering{}{\footnotesize{}\label{fig:4}}{\footnotesize\par}
\end{figure}

To better understand the proposed method in this paper, we investigate
behaviors of three approaches, two covariate adjusted and variable
selection approaches, through the Head start example where the approaches
are the same as those considered in the simulation, and Tables 5 and
6. The approach denoted by ``Cov-adjusted 1'' corresponds to the covariate
adjusted approach using the CCT MSE-optimal bandwidths, and the one
denoted by ``Cov-adjusted 2'' to the one using the CCFT MSE-optimal
bandwidths. We construct 50 subsamples of $k$ observations out of
the Head start data where $k$ is set to 1000, and 500. We estimate
the RDD treatment effects based on the subsamples using nine or forty-five
covariates where nine covariates are those considered in CCFT and
forty-five are those considered in Table 5. Figures \ref{fig:1}-\ref{fig:4}
show boxplots of the results on estimation (left panel) and inference
(right panel).\footnote{The results shown in Figures \ref{fig:1}-\ref{fig:3} are based on
the bandwidths with $h_{n}/b_{n}$ unrestricted, and those with the
bandwidths with $h_{n}/b_{n}=1$ are qualitatively similar.} The top and bottom of the boxes are the third and first quartiles,
and the top- and bottom-bars show the maximum and minimum values less
than (the third quartiles + 1.5 times the interquartile range) and
greater than (the first quartile - 1.5 times the interquartile range),
respectively. The left panels in Figures \ref{fig:1}-\ref{fig:4}
show the differences in the RDD treatment effects between the three
approaches and the standard one where the standard one is based on
CCT. The right panel in each figure shows the corresponding differences
in the confidence interval lengths.\footnote{The variable selection approach produced the identical result to the
standard one about 20 times for all setups. The boxplots are drawn
based on the non-identical results.}

First, the left panels of Figures \ref{fig:1} and \ref{fig:2} show
the point estimates by the covariate adjusted approaches deviate from
the standard one by a large extent when the number of covariates is
large, and the differences get larger as the sample size becomes small.
We observe a similar tendency for the situations with nine covariates,
while differences are less dramatic. Second, we note that the point
estimates based on the standard and variable selection approaches
are very close, and they are stable. Third, we observe that the covariate
adjusted approaches tend to produce large decreases in the interval
length, which can reflect under-coverages as observed in Table \ref{tab:inf}.
Fourth, the variable selection approach leads to reductions of around
5\%, which are still nonnegligible. Fifth, the results for nine covariates
given in Figures \ref{fig:3} and \ref{fig:4} are similar to those
for forty-five covariates, although the magnitude of changes for nine
covariates is not as extreme as that for forty-five covariates. These
results show the usefulness of our variable selection approach not
only for the situation where the number of covariates is large but
also for that where the number of covariates is relatively small,
when the number of observations is not so large.

We often encounter situations where a number of potentially useful
covariates are available. Furthermore, It is common to employ transformations
of covariates such as interaction and quadratic terms. However, we
are typically uncertain about which covariates contribute to improving
efficiency possibly due to the lack of economic theories. The RDD
analysis is local in nature, and the number of covariates relative
to the effective sample size can be pretty large. We have seen that
covariate adjusted approaches can become misleading in several examples.
The steady performance of the variable selection approach under various
circumstances is noteworthy, and it can provide an essential second
opinion for the standard and covariate adjusted approaches.